Papers with Cohen’s Kappa
A Multi-Axis Annotation Scheme for Event Temporal Relations (P18-1)
Copied to clipboard
| Challenge: | Existing temporal relation (TempRel) annotation schemes have low inter-annotator agreements even between experts, suggesting that the current annotation task needs a better definition. |
| Approach: | They propose to annotate temporal relation (TempRel) annotation schemes based on event start-points instead of a conventional 60’s-80’s model. |
| Outcome: | The proposed model improves IAA from the conventional 60’s to 80’s and can be used by crowdsourcing to alleviate labor intensity. |
ROSE: An Intent-Centered Evaluation Metric for NL2SQL (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics for evaluating the effectiveness of Natural Language to SQL (NL2SQL) solutions are becoming unreliable due to its sensitiveness to syntactic variation and inconsistent consistency with ground-truth SQL. |
| Approach: | They propose an intent-centered metric that focuses on whether the predicted SQL answers the question, rather than consistency with the ground-truth SQL. |
| Outcome: | The proposed metric outperforms the next-best metric by nearly 24% on the expert-aligned validation set **ROSE-VEC**. |
Establishing Annotation Quality in Multi-label Annotations (2022.coling-1)
Copied to clipboard
| Challenge: | Multi-label annotations allow multiple interpretations of a single item, but they also affect the chance that two coders agree with each other. |
| Approach: | They propose a bootstrapped method to obtain chance agreement for each measure and a method to get an adjusted agreement coefficient that is more interpretable. |
| Outcome: | The proposed method allows for an adjusted agreement coefficient that is more interpretable on simulated datasets. |
Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | In second language learning, summaries are among the most popular type of student assignments. |
| Approach: | They propose to revise the annotation guidelines to allow machine implementation of the new annotation guidelines. |
| Outcome: | The proposed algorithm achieves 0.789 precision and 0.844 recall over the L2WS 2021 corpus. |
Common Law Annotations: Investigating the Stability of Dialog System Output Annotations (2023.findings-acl)
Copied to clipboard
Seunggun Lee, Alexandra DeLucia, Nikita Nangia, Praneeth Ganedi, Ryan Guan, Rubing Li, Britney Ngaw, Aditya Singhal, Shalaka Vaidya, Zijun Yuan, Lining Zhang, João Sedoc
| Challenge: | High agreement is often used to show reliability of annotation procedures, but it is insufficient to ensure or reproducibility. |
| Approach: | They propose a protocol that increases Inter-Annotator Agreement among annotators and a standardized and codified protocol that strictly enforces transparency in the annotation process. |
| Outcome: | The proposed protocol ensures transparency in the annotation process, which ensures reproducibility of annotation guidelines. |